Papers with automatic classification
Detecting Primary Progressive Aphasia (PPA) from Text: A Benchmarking Study (2026.findings-eacl)
Copied to clipboard
Ghofrane Merhbene, Fabian Lecron, Philippe Fortemps, Bradford C. Dickerson, Mascha Kurpicz-Briki, Neguine Rezaii
| Challenge: | Primary progressive aphasia (PPA) is a neurodegenerative disorder characterized by progressive language deficits as the primary symptom. |
| Approach: | They benchmarked the performance of traditional machine learning models with various feature extraction techniques, transformer-based models, and large language models (LLMs) they found that transformer-Based models exceeded chance-level performance in terms of balanced accuracy, while MLP using MentalBert’s embeddings achieved the highest accuracy. |
| Outcome: | The proposed models outperform chance-level models in terms of balanced accuracy while using MentalBert’s embeddings achieve the highest accuracy. |
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)
Copied to clipboard
| Challenge: | Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts. |
| Approach: | They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues. |
| Outcome: | The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks. |
Multimodal Pipeline for Collection of Misinformation Data from Telegram (2022.lrec-1)
Copied to clipboard
| Challenge: | a large portion of misinformation is spread via multimodal means, such as images and videos . a new pipeline for collecting misinformation from Telegram allows us to collect a greater variety of mis-information examples . |
| Approach: | They propose to use AI to understand misinformation flow across social media platforms . they collect data from Telegram groups which promote COVID-19 misinformation . |
| Outcome: | The proposed dataset contains almost one million messages from 2k different public channels related to spreading COVID-19 misinformation. |
The Automatic Annotation of the Semiotic Type of Hand Gestures in Obama’ s Humorous Speeches (L18-1)
Copied to clipboard
| Challenge: | Existing studies on hand gestures from video-recorded speeches have not identified them. |
| Approach: | They annotated and analysed hand gestures produced by Barack Obama . they trained machine learning algorithms to classify the semiotic type of hand gesture . |
| Outcome: | The proposed method can be used to classify hand gestures on video-recorded speeches and in advanced multimodal interactive systems. |
Identifying the Human Values behind Arguments (2022.acl-long)
Copied to clipboard
| Challenge: | et al., 2003) examines human values in natural language arguments . authors provide a dataset of 5270 arguments from four geographical cultures . |
| Approach: | They propose a multi-level taxonomy of human values with 54 values and a dataset of 5270 arguments from four geographical cultures, manually annotated for human values. |
| Outcome: | The proposed model shows that human values are more diverse than previously thought . it shows that people disagree on the best course forward on controversial issues . |
Classifying Referential and Non-referential It Using Gaze (D18-1)
Copied to clipboard
| Challenge: | a particular problem for anaphora resolution systems is the pronoun it, which can be used both referentially and non-referentially. |
| Approach: | They use eye-tracking data to learn how humans perform disambiguation and use it to improve automatic classification. |
| Outcome: | The proposed system outperforms a baseline and outperformed linguistic-based approaches. |
Classification of Closely Related Sub-dialects of Arabic Using Support-Vector Machines (L18-1)
Copied to clipboard
| Challenge: | Existing studies on dialect identification have focused on binary classifications between colloquial Arabic and dialectal Egyptian . |
| Approach: | They propose to use an n-gram based SVM to classify on a fine-grained sub-dialectal level and compare it to methods used in dialect classification such as vocabulary pruning. |
| Outcome: | The proposed method is compared to methods used in dialect classification such as vocabulary pruning of shared items across dialects. |
A Corpus for Suggestion Mining of German Peer Feedback (2022.lrec-1)
Copied to clipboard
| Challenge: | e.g. Massive Open Online Courses (MOOCs) are increasingly important to meet the demand for feedback in large scale classes. |
| Approach: | They propose to use peer feedback to detect suggestions on how to improve the work of students in a german university course. |
| Outcome: | The proposed corpus is the first student peer feedback corpus in germany and has been labelled with a new annotation scheme. |
Hierarchical Multi-Label Classification of Scientific Documents (2022.emnlp-main)
Copied to clipboard
| Challenge: | Automated topic classification is a useful tool for managing scientific documents in a digital collection. |
| Approach: | They propose a hierarchical multi-label text classification dataset with keyword labeling as an auxiliary task. |
| Outcome: | The proposed model achieves a Macro-F1 score of 34.57% and is publicly available. |
Invisible to People but not to Machines: Evaluation of Style-aware HeadlineGeneration in Absence of Reliable Human Judgment (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a data alignment strategy and different training/testing settings, we aim at decoupling content from style and preserving the latter in generation. |
| Approach: | They propose a fine-grained evaluation strategy based on automatic classification to evaluate generated headlines' quality in terms of their newspaper-compliance. |
| Outcome: | The proposed model learns newspaper-specific style, but humans aren't reliable judges for this task, and deserves particular care in its design. |
Murre24: Dialect Identification of Finnish Internet Forum Messages (2024.lrec-main)
Copied to clipboard
| Challenge: | 94 million messages posted on the largest Finnish internet forum, Suomi24, are classified to present either the standard language, one of the seven traditional dialects, a colloquial style or the Helsinki slang. |
| Approach: | They present a collection of dialectal messages posted on the largest Finnish internet forum, Suomi24 . they manually annotated a dataset and used it to train dialect identification models . |
| Outcome: | The proposed method is the best for differentiating standard Finnish from non-standard Finnish, while fine-tuning a BERT-based model achieves best scores on the final dialect identification task. |